Papers with pretrained word embeddings

13 papers
VSP at PharmaCoNER 2019: Recognition of Pharmacological Substances, Compounds and Proteins with Recurrent Neural Networks in Spanish Clinical Cases (D19-57)

Copied to clipboard

Challenge: The Named Entity Recognition of drugs, medications and chemical entities in Spanish is a new task in the field of NLP .
Approach: They propose to use SNOMED CT term search engine to classify the entities in Spanish and a neural model for the Named Entity Recognition.
Outcome: The proposed system achieves 76.29% and 60.34% performance in the Named Entity Recognition and Concept indexing tasks.
Transfer Learning in Natural Language Processing (N19-5)

Copied to clipboard

Challenge: supervised machine learning is based on learning in isolation, a single predictive model for a task using a dataset.
Approach: They present an overview of modern transfer learning methods in natural language processing . they review examples and case studies on how models can be integrated and adapted .
Outcome: The proposed methods improve upon the state-of-the-art on a wide range of NLP tasks.
Multimodal Machine Translation with Embedding Prediction (N19-3)

Copied to clipboard

Challenge: Pretrained word embeddings improve multimodal machine translation of low-resource domains due to a shortage of training data.
Approach: They propose to combine pretrained word embeddings with search-based approaches to improve NMT of low-resource domains to better translate rare words.
Outcome: The proposed approach improves translation performance by 1.24 METEOR and 2.49 BLEU and achieves 7.67 F-score.
Comparing Pretrained Multilingual Word Embeddings on an Ontology Alignment Task (L18-1)

Copied to clipboard

Challenge: Existing word embeddings capture a string's semantics and can be trained for multiple languages.
Approach: They propose to compare three different multilingual pretrained word embedding repositories with a string-matching baseline and use it to compute semantic similarities of strings in different languages.
Outcome: The proposed method produces correct alignments on a non-standard dataset on all four languages.
Extracting Possessions from Social Media: Images Complement Language (D19-1)

Copied to clipboard

Challenge: Existing studies show that authors of tweets possess objects they tweet about.
Approach: They propose a dataset and experiments to determine whether tweet authors possess objects they tweet about.
Outcome: The proposed strategy incorporates visual information into any neural network beyond weights from pretrained networks.
Adaptation of Hierarchical Structured Models for Speech Act Recognition in Asynchronous Conversation (N19-1)

Copied to clipboard

Challenge: asynchronous domains lack large labeled datasets to train an effective speech act recognition model.
Approach: They propose methods to leverage abundant unlabeled conversational data and available labeled data from synchronous domains to train an effective SAR model.
Outcome: The proposed method outperforms existing methods when trained on in-domain data only.
Incorporating Emoji Descriptions Improves Tweet Classification (N19-1)

Copied to clipboard

Challenge: Tweets are short messages that often include specialized language such as hashtags and emojis.
Approach: They propose a simple strategy to replace emojis with their natural language description and use pretrained word embeddings to process tweets.
Outcome: The proposed method is more effective than pretrained emoji embeddings for tweet classification.
Early Discovery of Disappearing Entities in Microblogs (2023.acl-long)

Copied to clipboard

Challenge: a study on detecting disappearing entities from noisy microblogs has been published on the real world . a major challenge is detecting uncertain contexts of disappearing entity from noisy posts .
Approach: They propose to use Twitter to detect disappearing entities from noisy microblogs . they build large-scale Twitter datasets of disappearing entity and refine word embeddings based on these data .
Outcome: The proposed method outperforms baseline methods on noisy microblog streams and more than 70% of disappearing entities in Wikipedia are discovered earlier than the update on Wikipedia.
Probing the Probing Paradigm: Does Probing Accuracy Entail Task Relevance? (2021.eacl-main)

Copied to clipboard

Challenge: Neural models have established state-of-the-art performance on several NLP benchmarks, but little is understood about the mechanisms by which they operate.
Approach: They examine the probing paradigm through a set of controlled synthetic tasks and show that pretrained word embeddings play a considerable role in encoding these properties rather than the training task itself.
Outcome: The proposed model can encode linguistic properties above chance-level even when distributed in the data as random noise, reversing the interpretation of absolute claims on probing tasks.
Multi-source Neural Topic Modeling in Multi-view Embedding Spaces (2021.naacl-main)

Copied to clipboard

Challenge: Recent work has used pre-trained word embeddings to address data sparsity in short-text or small document collections.
Approach: They propose a neural topic modeling framework using multi-view embedding spaces to improve topic quality and deal with polysemy.
Outcome: The proposed framework improves topic quality and deal with polysemy.
Delta-training: Simple Semi-Supervised Text Classification using Pretrained Word Embeddings (D19-1)

Copied to clipboard

Challenge: Pretrained word embeddings outperforms classifiers with randomly initialized word embeds, a new method is proposed for semi-supervised text classification.
Approach: They propose a method that uses pretrained word embeddings to predict text classification . they use unlabeled data to build a classifier, and use early-stopping to improve performance .
Outcome: The proposed method outperforms self-training and co-training frameworks on unlabeled data.
An Empirical Study on Leveraging Position Embeddings for Target-oriented Opinion Words Extraction (2021.emnlp-main)

Copied to clipboard

Challenge: Current methods for extracting opinion words for an aspect in text leverage position embeddings to capture relative position of word to the target.
Approach: They propose to use pretrained word embeddings to extract opinion words for a given aspect in text.
Outcome: The proposed methods outperform current methods on a task based on pre-trained word embeddings and position embedders.
Revisiting Tri-training of Dependency Parsers (2021.emnlp-main)

Copied to clipboard

Challenge: Pre-trained word embeddings and self-training have been used in dependency parsing tasks for years.
Approach: They compare tri-training and pretrained word embeddings in dependency parsing . they use language-specific FastText and ELMo embedds and multilingual BERT embedders .
Outcome: The proposed methods are tri-training and pretrained word embeddings.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations